Statistical analysis of filled pauses2 rhythm for disfluent speech synthesis
نویسندگان
چکیده
Given that state of the art speech synthesis systems have already reached a high naturalness level, it is time to move to talking speech from the actual read speech framework. For this purpose it is thus necessary to investigate how disfluencies can be included in speech synthesis and even increase its naturalness. This paper builds on a previously presented work and focuses on finding a local model of filled pauses rhythm. A statistical study of rhythm effects around filled pauses is presented and based on the correlation between rhythm variables, a regression model is proposed to predict filled pauses duration and prepausal lengthening.
منابع مشابه
Statistical analysis of filled pauses’ rhythm for disfluent speech synthesis
Given that state of the art speech synthesis systems have already reached a high naturalness level, it is time to move to talking speech from the actual read speech framework. For this purpose it is thus necessary to investigate how disfluencies can be included in speech synthesis and even increase its naturalness. This paper builds on a previously presented work and focuses on finding a local ...
متن کاملDisfluent Speech Analysis and Synthesis: a preliminary approach
Despite of the existence of high quality unit selection speech synthesizers, they are based on a reading style approach. However, new applications such as Speech-to-Speech Translation or Speech User Interfaces demand a talking style which is more natural in these contexts. Disfluencies are a major characteristic of talking style so that it is convenient to be able to generate disfluent speech. ...
متن کاملBreath and Non-breath Pauses in Fluent and Disfluent Phases of German and French L1 and L2 Read Speech
In this study we examined the read speech of native and nonnative speakers with respect to pausing details of audible breathing, particularly in disfluent phases. 20 German and 20 French native speakers read the same narrative text in their native (L1) and in their non-native language (L2). Some expected results were confirmed: more frequent pauses and more frequent disfluencies in L2, as well ...
متن کاملOn the generation of synthetic disfluent speech: local prosodic modifications caused by the insertion of editing terms
Disfluent speech synthesis is necessary in some applications such as automatic film dubbing or spoken translation. This paper presents a model for the generation of synthetic disfluent speech based on inserting each element of a disfluency in a context where they can be considered fluent. Prosody obtained by the application of standard techniques on these new sentences is used for the synthesis...
متن کاملA Lattice-based Approach to Automatic Filled Pause Insertion
This paper describes a novel method for automatically inserting filled pauses (e.g., UM) into fluent texts. Although filled pauses are known to serve a wide range of psychological and structural functions in conversational speech, they have not traditionally been modelled overtly by state-of-the-art speech synthesis systems. However, several recent systems have started to model disfluencies spe...
متن کامل